免费领取大会全套演讲PPT    

点击领取

我要参会

Shuwen Wang

Core Contributor to SGLang, R&D Engineer at RadixArk

Shuwen Wang is a core contributor to SGLang and a maintainer of HiCache at RadixArk. He is primarily involved in the development and optimization of HiCache, Unified Radix Cache, and speculative decoding. His work focuses on KV cache management, hierarchical caching, and inference performance optimization for large language models. He also contributes to code review and maintenance within the open-source community.

Topic

Context Caching for Agentic Inference: From RadixAttention to Unified Radix Cache

This talk will introduce the evolution of context caching in SGLang, from RadixAttention to Unified Radix Cache, and explore its design for Agentic Inference, hybrid models, and multi-level caching scenarios. Outline RadixAttention and Prefix Caching New Challenges in Agentic Inference Design and Implementation of Unified Radix Cache HiCache and Future Evolution Key Takeaways Gain an understanding of the core design and evolution of SGLang’s KV Cache, the motivations behind these architectural changes, and the key challenges of context reuse and cache management in Agentic Inference scenarios.

© boolan.com 博览 版权所有

沪ICP备15014563号-6

沪公网安备31011502003949号