Chen
50 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
This paper introduces X-Stage, a software-visible post-issue pipeline stage to improve communication efficiency in distributed diffusion transformer (DiT) inference, leading to significant speedups for DeepGEMM MegaMoE and Ulysses sequence-parallel attention.
Libra is a load balancing approach for long-context LLM training that groups packed sequences into fixed-size pools and reduces attention workload variance, improving end-to-end throughput and straggler-attention speedup.
This paper proposes Gleam, a framework for efficient GPU sharing across local-area CUDA devices, reducing bandwidth overhead, improving API call latency, and ensuring context consistency.
The paper provides a negative answer to Carlson's question about whether the depth of a finite-group cohomology ring is always realized by the dimension of one of its associated primes.
This paper presents TRUAV, a distributed multi-agent reinforcement learning framework for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs, eliminating the need for global state exchange.
This paper benchmarks zero-shot synthesis of parent-selection operators across eight large language models and finds that Claude Sonnet~4.6 and Gemini~3.1 Pro perform strongly, with the best operator surpassing automatic baselines.
The paper proposes TRIDENT, a framework to restore a source speaker's identity from converted audio using a three-pronged architecture.
This paper introduces ClinFusion, a vision-centric multimodal large language model designed for holistic medical understanding, featuring a Cascade Spatial-Aware Locality Fusion operator and a vision-grounded evaluation framework.
This paper organizes embodied data sources for multimodal foundation models into a pyramid, focusing on real-robot, UMI-style, egocentric and exocentric, simulation, and general vision-language data.
This paper presents DOPS, a hardware-aware framework for optimizing operator scheduling and weight layouts in Large Language Models, achieving significant speedups over prefill-decode disaggregation.
This paper proposes CoSMIC, a framework for semantic integrated sensing and communication (ISAC) that embeds semantic information into waveforms while maintaining covertness and sensing fidelity.
This paper introduces Desktop-Delta Bench (DDB), an offline step-level benchmark for evaluating computer-use agents' ability to reconstruct causal transitions in desktop GUI environments.
This paper introduces RSIBench-Data, a controlled benchmark for evaluating data-centric research capabilities of LLM agents.
This paper introduces CW-Ghost, a method for estimating cache line fill volume and determining helper-thread prefetching granularity based on cache capacity constraints.
Specula is an autonomous system that generates high-quality formal specifications for large, complex code using LLMs, improving understanding and finding bugs.
The paper proposes SharpRec, a framework for LLM-based Cross-Domain Sequential Recommendation to address the bottlenecks of cross-domain knowledge conflict and performance saturation in multi-domain fusion.
This paper adapts Large Language Models as semantic representation backbones in a two-tower retrieval architecture for high-throughput, large-scale recommendation systems.
This paper presents a compositional cost analysis for probabilistic programs with hierarchical cost structures, allowing computation of mean and higher moments of non-additive costs.
This paper proposes a randomized iterative method called Sequential Preconditioned Conjugate Gradient Method (SPCG) for large-scale linear statistical models, which significantly reduces computational cost by solving smaller subproblems.
This paper reveals a new vulnerability in Language Model (LLM) inference efficiency caused by persona consistency and proposes RolePlay, a framework to amplify inference costs.
Papers
Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling
Jiaqi Yang, Jiayi Li, Yihan Fu, Hongxiao Zhao +4 more
This paper presents DOPS, a hardware-aware framework for optimizing operator scheduling and weight layouts in Large Language Models, achieving significant speedups over prefill-decode disaggregation.