Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Home/Authors/Tianyu Wo

Tianyu Wo

2 indexed papers

Recent (6 mo)
2
With code
0
Influential cites
0
Benchmarked
0

Publications per year

2
26

Top categories

Distributed×2AI×2ML×2Performance×2

Frequent co-authors

Chunming Hu2×
Renyu Yang2×
Yihui Zhang1×
Jinghao Wang1×
Xiaoyang Sun1×
Menghao Zhang1×

Research Timeline

2026
CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

This paper proposes CrossPool, a serving engine for cold Machine Learning Models (LLMs) that separates weights and KV-cache into two GPU memory pools to improve GPU memory utilization and long-context support.

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

This paper presents SpecBox, a runtime system for LLM agents that uses speculative sandbox preallocation to improve resource utilization and reduce interactive tail latency.

Highlighted terms show continued research focus across papers

Papers

cs.DCcs.AIcs.LGEmpiricalRecentJul 27, 2026

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

Yihui Zhang, Tianyu Wo, Jinghao Wang, Xiaoyang Sun +6 more

This paper presents SpecBox, a runtime system for LLM agents that uses speculative sandbox preallocation to improve resource utilization and reduce interactive tail latency.

View →
cs.DCcs.AIcs.LGEmpirical
Recent
Jun 23, 2026

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Zhuoren Ye, Tianyu Wo, Dinghao Xue, Mingming Zhang +3 more

This paper proposes CrossPool, a serving engine for cold Machine Learning Models (LLMs) that separates weights and KV-cache into two GPU memory pools to improve GPU memory utilization and long-context…

View →