Shaohuai Shi
2 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
DigenRL is a disaggregated RL framework for diffusion-based generative LLMs that achieves 1.56-2.10x throughput improvements over state-of-the-art diffusion RL systems.
KernelFlume is a decode-centric architecture that disaggregates the stable projection/FFN path from core-attention computation to improve efficiency and reduce cost in serving long-context demand.
Papers
KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding
Guangyu Xiang, Xueze Kang, Lin Zhang, Wenxiang Lin +3 more
KernelFlume is a decode-centric architecture that disaggregates the stable projection/FFN path from core-attention computation to improve efficiency and reduce cost in serving long-context demand.