Shuwen Wang
Core Contributor to SGLang, R&D Engineer at RadixArk
Shuwen Wang is a core contributor to SGLang and a maintainer of HiCache at RadixArk. He is primarily involved in the development and optimization of HiCache, Unified Radix Cache, and speculative decoding. His work focuses on KV cache management, hierarchical caching, and inference performance optimization for large language models. He also contributes to code review and maintenance within the open-source community.
Topic
Context Caching for Agentic Inference: From RadixAttention to Unified Radix Cache
This talk will introduce the evolution of context caching in SGLang, from RadixAttention to Unified Radix Cache, and explore its design for Agentic Inference, hybrid models, and multi-level caching scenarios. Outline RadixAttention and Prefix Caching New Challenges in Agentic Inference Design and Implementation of Unified Radix Cache HiCache and Future Evolution Key Takeaways Gain an understanding of the core design and evolution of SGLang’s KV Cache, the motivations behind these architectural changes, and the key challenges of context reuse and cache management in Agentic Inference scenarios.