Zishen Wan
2 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper presents Dyserve, a workflow-aware serving layer for agentic AI applications that compiles per-node model and verifier choices into an integer linear program, allowing for efficient model selection and verification under serving-time conditions.
This paper introduces ArchEval, a benchmark and platform for evaluating LLM agents on computer architecture design and optimization.
Papers
A Workflow-Aware Serving Layer for Agentic Applications
Jiayi Qian, Zishen Wan, Hanchen Yang, Chun Tao +2 more
The paper presents Dyserve, a workflow-aware serving layer for agentic AI applications that compiles per-node model and verifier choices into an integer linear program, allowing for efficient model se…