~ similar to 2606.23945· 20 results
Yilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao +3 more
The paper introduces COSM, a cooperative scheduling framework to facilitate concurrent operation of Processing-in-Memory (PIM) and CPU tasks on mobile platforms, improving PIM throughput by up to 2.8x…
This paper argues that idle cores on contemporary multicore processors can return compute capacity and proposes a budget-centric view of service core systems.
The I/O Resource Manager (IORM) is presented as a multi-stage distributed scheduler to maintain consistent performance and enforce global I/O limits in shared-nothing disaggregated storage clusters.
Junwei Ji, Woon-Seng Gan, Boxiang Wang, Ziyi Yang +1 more
This paper proposes an adaptive momentum term for the ASSS-MGDFxLMS algorithm in distributed multichannel active noise control systems to accelerate convergence while maintaining robustness under comm…
Hai Duc Nguyen, Bogdan Nicolae, Tekin Bicer, Amal Gueroudji +3 more
This paper presents two techniques, dynamic checkpointing and progress-aware load redistribution, to maintain forward progress and balanced execution in real-time scientific workflows using the produc…
Valinor is a hardware-OS cooperative memory allocation substrate that introduces a programmable hardware allocation engine for improving performance and reducing energy consumption in virtual-to-physi…
This paper proposes a dependency-aware autoscaling framework for serverless computing, integrating graph-based bottleneck identification, short-term workload forecasting, multi-model consensus, and co…
Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin +5 more
Proposed EnerInfer framework manages energy efficiency, throughput, and thermal comfort for on-device LLM inference, improving energy efficiency up to 65% without QoE violation.
The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…
F. Nisa Bostanci, Haocong Luo, Ataberk Olgun, Maria Makeenkova +3 more
The authors of Ramulator 2.0 simulator challenge the claims made in a research paper about its performance and propose best practices to avoid simulator usage errors.
Hyunwoo Oh, Suyeon Jang, Hanning Chen, Sanggeon Yun +2 more
The paper presents ExaGEMM, a framework for designing and exploring CPU-native low-bit GEMM via register-resident LUT execution.
Zhuoren Ye, Tianyu Wo, Dinghao Xue, Mingming Zhang +3 more
This paper proposes CrossPool, a serving engine for cold Machine Learning Models (LLMs) that separates weights and KV-cache into two GPU memory pools to improve GPU memory utilization and long-context…
This paper analyzes the synchronization of power usage in large-scale AI training facilities and identifies the coupling channel in load-dependent throttling.
Aquifer is the first system to serve MicroVM snapshots from a hierarchical CXL+RDMA memory pool, achieving 2.2x geometric-mean speedup in end-to-end invocation time over Firecracker.
The paper introduces Chimera, a highly efficient and scalable MCU designed for ultra-low-power edge AI inference, achieving 3.1 TOPS/W by integrating a dedicated transformer accelerator and a QoS-guar…
This paper proposes CAPS, a scheduling layer for data centers that separates rate computation and packet scheduling, reducing queue occupancy by up to 10x without throughput loss.
The paper proposes a method to improve the performance of fine-grained offloads on servers by overlapping the offload with other requests using server-side routing.