ArXivCSExplorer
β˜†β˜†BookmarksπŸ†RSSHow to UseFAQ
Built with and by Teycir Ben Soltaneβ€’
How to Useβ€’FAQβ€’GitHubβ€’arXiv.orgβ€’
Share:

20 results for β€œworkload inference”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.β“˜

Want pure semantic search? Try claim verification β†’

cs.ARcs.LGEmpiricalRecentJun 11, 2026

BigPower: Hierarchical Source-Level Module Power Estimation for CPUs with Large Language Models

Honghua Zhu, Chunjie Luo, Jianfeng Zhan

This paper introduces BigPower, a hierarchical source-level surrogate model for fine-grained module-level power estimation during CPU design using large language models and architectural hierarchy.

View β†’
cs.ARcs.LGcs.OSEmpiricalRecentJul 16, 2026

ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM

Hyunwoo Oh, Suyeon Jang, Hanning Chen, Sanggeon Yun +2 more

The paper presents ExaGEMM, a framework for designing and exploring CPU-native low-bit GEMM via register-resident LUT execution.

View β†’
cs.NIEmpiricalRecentJul 28, 2026

Round Trip Time: A Benign Signal or an Indirect Window into Datacenter Workloads?

Sourya Saha, Md Nurul Absur, Saptarshi Debroy

This paper investigates a network side-channel vulnerability in multi-tenant datacenter fabrics caused by shared congestion behavior, achieving up to 97.3% run-level accuracy in workload inference.

View β†’
cs.AREmpiricalRecentJul 28, 2026

Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling

Jiaqi Yang, Jiayi Li, Yihan Fu, Hongxiao Zhao +4 more

This paper presents DOPS, a hardware-aware framework for optimizing operator scheduling and weight layouts in Large Language Models, achieving significant speedups over prefill-decode disaggregation.

View β†’
cs.AReess.SYEmpiricalRecentJul 24, 2026

Reducing Instruction-Fetch Energy in RISC-V for Embedded AI Processing via Dynamic and Static Loop Caching

Wiebren Wijnstra, Sameed Sohail, Berend-Jan van der Zwaag, Sabih Gerez +1 more

This paper presents two loop cache architectures for RISC-V processors to reduce instruction fetches and energy consumption during AI inference.

View β†’
cs.LGcs.AIRecentMay 31, 2026

Beyond Task-Agnostic: Task-Aware Grouping for Communication-Efficient Multi-Task MoE Inference

Zhiyao Xu, Aoxue Liu, Zhanjie Ding, Dan Zhao +2 more

The paper proposes Task-Aware Coactivation Grouping (TACG) to significantly reduce communication costs in multi-task MoE inference by grouping experts based on task-specific co-activation patterns, ou…

View β†’
cs.PFcs.AREmpiricalRecentJul 16, 2026

Campaign Diagrams: Visualizing the March Through the Phases of a Workload

Toluwanimi O. Odemuyiwa, John D. Owens, Michael Pellauer, Joel S. Emer

This paper introduces campaign diagrams, a visualization technique for analyzing resource utilization and identifying bottlenecks in modern workloads.

View β†’
cs.CRcs.SEeess.SPRecentApr 11, 2026

Organizational Security Resource Estimation via Vulnerability Queueing

Abdullah Y. Etcibasi, Zachary Dobos, C. Emre Koksal

The paper proposes a dynamic queueing framework that estimates an organization's cyber resources and attack surface dynamics by analyzing the timestamps of vulnerabilities and fixes, achieving high ac…

View β†’
cs.DCcs.MATheoreticalRecentJul 3, 2026

A Workflow-Aware Serving Layer for Agentic Applications

Jiayi Qian, Zishen Wan, Hanchen Yang, Chun Tao +2 more

The paper presents Dyserve, a workflow-aware serving layer for agentic AI applications that compiles per-node model and verifier choices into an integer linear program, allowing for efficient model se…

View β†’
cs.DCcs.SENEWEmpiricalJul 29, 2026

Hybrid Workflow Composition for Extreme-Scale Data Processing: A Case Study on the HL-LHC (Extended Version)

Alan Malta Rodrigues, Douglas Thain

This paper presents a simulation framework to optimize workflow composition in high-throughput computing environments, demonstrating up to 3.8x throughput increase and a 14.9x reduction in network ove…

View β†’
cs.IRcs.AIcs.LGEmpiricalRecentJun 26, 2026

End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference

Yuhang Chen, Jinhao Duan, Ruichen Zhang, Mingfu Liang +10 more

This paper proposes Learning to Allocate (L2A), an end-to-end framework for resource-adaptive inference in Large Language Models (LLMs) using budget-conditioned and input-aware gating networks.

View β†’
cs.ARcs.AIEmpiricalRecentJul 6, 2026

Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference

Guanyu Cai, Ruiming Tian, Lang Yang, Zhouhong Ren +3 more

This paper presents a comprehensive measurement study on the performance and energy consumption of large language models on mobile devices, using five frameworks and three hardware backends, and intro…

View β†’
cs.DCEmpiricalRecentJul 12, 2026

Ichnos+: Estimating the Carbon Footprint of Scientific Workflows Using Fitted Power Models

Kathleen West, Youssef Moawad, Philipp Thamm, Vasilis Bountris +4 more

This paper introduces Ichnos+, a system to estimate the environmental footprint of Nextflow scientific workflows using post-hoc analysis and node-specific power models.

View β†’
cs.SEcs.LGcs.OSEmpiricalRecentJun 22, 2026

EnerInfer: Energy-Aware On-Device LLM Inference

Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin +5 more

Proposed EnerInfer framework manages energy efficiency, throughput, and thermal comfort for on-device LLM inference, improving energy efficiency up to 65% without QoE violation.

View β†’
cs.NIcs.DCcs.LGEmpiricalRecentJul 28, 2026

Incast-Free MoE Rate-Based Scheduling

Evyatar Cohen, Jose Yallouz, Alexander Shpiner, Mark Silberstein +2 more

This paper proposes a proactive fair scheduling framework to prevent fabric oversubscription and eliminate incast in Mixture of Experts (MoE) architectures, demonstrating consistent link utilization a…

View β†’
cs.DCEmpiricalRecentJul 28, 2026

CW-Ghost: Search-Free Granularity Selection for Helper-Thread Prefetching via Capacity Windows

Ya Zhang, Tong Lei, Yao Chen, Yonggang Che +3 more

This paper introduces CW-Ghost, a method for estimating cache line fill volume and determining helper-thread prefetching granularity based on cache capacity constraints.

View β†’
cs.CRcs.AIcs.LGRecentMay 11, 2026

Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing

Yunze Zhao, Yibo Zhao, Yuchen Zhang, Zaoxing Liu +1 more

The paper introduces GRIEF, a greybox fuzzer that discovers critical, concurrency-related vulnerabilities in LLM serving systems by treating timed multi-request traces as inputs, finding issues like c…

View β†’