20 results for βworkload inferenceβ
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.β
Want pure semantic search? Try claim verification β
This paper introduces BigPower, a hierarchical source-level surrogate model for fine-grained module-level power estimation during CPU design using large language models and architectural hierarchy.
Hyunwoo Oh, Suyeon Jang, Hanning Chen, Sanggeon Yun +2 more
The paper presents ExaGEMM, a framework for designing and exploring CPU-native low-bit GEMM via register-resident LUT execution.
This paper investigates a network side-channel vulnerability in multi-tenant datacenter fabrics caused by shared congestion behavior, achieving up to 97.3% run-level accuracy in workload inference.
Jiaqi Yang, Jiayi Li, Yihan Fu, Hongxiao Zhao +4 more
This paper presents DOPS, a hardware-aware framework for optimizing operator scheduling and weight layouts in Large Language Models, achieving significant speedups over prefill-decode disaggregation.
This paper presents two loop cache architectures for RISC-V processors to reduce instruction fetches and energy consumption during AI inference.
Zhiyao Xu, Aoxue Liu, Zhanjie Ding, Dan Zhao +2 more
The paper proposes Task-Aware Coactivation Grouping (TACG) to significantly reduce communication costs in multi-task MoE inference by grouping experts based on task-specific co-activation patterns, ouβ¦
This paper introduces campaign diagrams, a visualization technique for analyzing resource utilization and identifying bottlenecks in modern workloads.
The paper proposes a dynamic queueing framework that estimates an organization's cyber resources and attack surface dynamics by analyzing the timestamps of vulnerabilities and fixes, achieving high acβ¦
Jiayi Qian, Zishen Wan, Hanchen Yang, Chun Tao +2 more
The paper presents Dyserve, a workflow-aware serving layer for agentic AI applications that compiles per-node model and verifier choices into an integer linear program, allowing for efficient model seβ¦
This paper presents a simulation framework to optimize workflow composition in high-throughput computing environments, demonstrating up to 3.8x throughput increase and a 14.9x reduction in network oveβ¦
Yuhang Chen, Jinhao Duan, Ruichen Zhang, Mingfu Liang +10 more
This paper proposes Learning to Allocate (L2A), an end-to-end framework for resource-adaptive inference in Large Language Models (LLMs) using budget-conditioned and input-aware gating networks.
Guanyu Cai, Ruiming Tian, Lang Yang, Zhouhong Ren +3 more
This paper presents a comprehensive measurement study on the performance and energy consumption of large language models on mobile devices, using five frameworks and three hardware backends, and introβ¦
Kathleen West, Youssef Moawad, Philipp Thamm, Vasilis Bountris +4 more
This paper introduces Ichnos+, a system to estimate the environmental footprint of Nextflow scientific workflows using post-hoc analysis and node-specific power models.
Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin +5 more
Proposed EnerInfer framework manages energy efficiency, throughput, and thermal comfort for on-device LLM inference, improving energy efficiency up to 65% without QoE violation.
This paper proposes a proactive fair scheduling framework to prevent fabric oversubscription and eliminate incast in Mixture of Experts (MoE) architectures, demonstrating consistent link utilization aβ¦
Ya Zhang, Tong Lei, Yao Chen, Yonggang Che +3 more
This paper introduces CW-Ghost, a method for estimating cache line fill volume and determining helper-thread prefetching granularity based on cache capacity constraints.
Yunze Zhao, Yibo Zhao, Yuchen Zhang, Zaoxing Liu +1 more
The paper introduces GRIEF, a greybox fuzzer that discovers critical, concurrency-related vulnerabilities in LLM serving systems by treating timed multi-request traces as inputs, finding issues like cβ¦