ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

~ similar to 2606.23945· 20 results

cs.ARcs.DCEmpiricalRecentJun 29, 2026

COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile Devices

Yilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao +3 more

The paper introduces COSM, a cooperative scheduling framework to facilitate concurrent operation of Processing-in-Memory (PIM) and CPU tasks on mobile platforms, improving PIM throughput by up to 2.8x…

View →
cs.OScs.ARcs.NIEmpiricalRecentJul 17, 2026

Rethinking Polling Efficiency in Service Core Network Stacks

Matheus Stolet, Simon Peter, Antoine Kaufmann

This paper argues that idle cores on contemporary multicore processors can return compute capacity and proposes a budget-centric view of service core systems.

View →
cs.DCcs.OScs.PFEmpiricalRecentJul 18, 2026

Hardware-Transparent I/O Governance in Disaggregated Heterogeneous Storage

Rajarshi Chowdhury, Akshay Shah, Sue K. Lee

The I/O Resource Manager (IORM) is presented as a multi-stage distributed scheduler to maintain consistent performance and enforce global I/O limits in shared-nothing disaggregated storage clusters.

View →
eess.ASeess.SPEmpiricalRecentJul 19, 2026

Adaptive Momentum Enhanced Distributed Multichannel Active Noise Control for Faster Convergence under Communication Delays

Junwei Ji, Woon-Seng Gan, Boxiang Wang, Ziyi Yang +1 more

This paper proposes an adaptive momentum term for the ASSS-MGDFxLMS algorithm in distributed multichannel active noise control systems to accelerate convergence while maintaining robustness under comm…

View →
cs.DCEmpiricalRecentJun 29, 2026

StreamGuard: Low-Overhead Resilience for Real-time HPC Data Streams

Hai Duc Nguyen, Bogdan Nicolae, Tekin Bicer, Amal Gueroudji +3 more

This paper presents two techniques, dynamic checkpointing and progress-aware load redistribution, to maintain forward progress and balanced execution in real-time scientific workflows using the produc…

View →
cs.AREmpiricalRecentJul 16, 2026

Valinor: Architectural Support for Fast, Energy-Efficient and Programmable Physical Memory Allocation

Konstantinos Kanellopoulos, Spiros Galanopoulos, Konstantinos Sgouras, Vlad-Petru Nitu +6 more

Valinor is a hardware-OS cooperative memory allocation substrate that introduces a programmable hardware allocation engine for improving performance and reducing energy consumption in virtual-to-physi…

View →
cs.LGcs.AIcs.DCEmpiricalRecentJul 16, 2026

An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism

Mobina Kashaniyan, Mehrdad Ashtiani, Amirhossein Ghassemi

This paper proposes a dependency-aware autoscaling framework for serverless computing, integrating graph-based bottleneck identification, short-term workload forecasting, multi-model consensus, and co…

View →
cs.SEcs.LGcs.OSEmpiricalRecentJun 22, 2026

EnerInfer: Energy-Aware On-Device LLM Inference

Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin +5 more

Proposed EnerInfer framework manages energy efficiency, throughput, and thermal comfort for on-device LLM inference, improving energy efficiency up to 65% without QoE violation.

View →
cs.PFcs.ARcs.DCRecentMay 27, 2026

Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory

Myeong Jun Jo

The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…

View →
cs.ARcs.PFEmpiricalRecentJun 12, 2026

Extended Abstract: Re-Evaluating the Real-System Modeling Accuracy of Ramulator 2.0

F. Nisa Bostanci, Haocong Luo, Ataberk Olgun, Maria Makeenkova +3 more

The authors of Ramulator 2.0 simulator challenge the claims made in a research paper about its performance and propose best practices to avoid simulator usage errors.

View →
cs.ARcs.LGcs.OSEmpiricalRecentJul 16, 2026

ExaGEMM: Exploration Framework for CPU-Driven ML Inference via Associative In-Register Computing for Low-Bit GEMM

Hyunwoo Oh, Suyeon Jang, Hanning Chen, Sanggeon Yun +2 more

The paper presents ExaGEMM, a framework for designing and exploring CPU-native low-bit GEMM via register-resident LUT execution.

View →
cs.DCcs.AIcs.LGEmpiricalRecentJun 23, 2026

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Zhuoren Ye, Tianyu Wo, Dinghao Xue, Mingming Zhang +3 more

This paper proposes CrossPool, a serving engine for cold Machine Learning Models (LLMs) that separates weights and KV-cache into two GPU memory pools to improve GPU memory utilization and long-context…

View →
eess.SYcs.DCnlin.AOTheoreticalRecentJul 22, 2026

Do Co-Located AI Training Jobs Synchronize? Load-Dependent Throttling as a Coupling Mechanism for Phase-Locking Behind a Shared Power Cap

Brieuc Le Roux Tardif

This paper analyzes the synchronization of power usage in large-scale AI training facilities and identifies the coupling channel in load-dependent throttling.

View →
cs.DCcs.OSEmpiricalRecentJun 23, 2026

Aquifer: Hierarchical Memory Pooling with CXL and RDMA for MicroVM Snapshots

Junliang Hu, Huaicheng Li, Ming-Chang Yang

Aquifer is the first system to serve MicroVM snapshots from a hierarchical CXL+RDMA memory pool, achieving 2.2x geometric-mean speedup in end-to-end invocation time over Firecracker.

View →
cs.ARRecentJun 1, 2026

CHIMERA: A Flexible and Scalable 3.1 TOPS/W AI-MCU with Transformer Accelerator and 563 Gb/s Shared-L2 Memory Subsystem with QoS Guarantees

Lorenzo Leone, Philip Wiese, Gamze İslamoğlu, Michael Rogenmoser +3 more

The paper introduces Chimera, a highly efficient and scalable MCU designed for ultra-low-power edge AI inference, achieving 3.1 TOPS/W by integrating a dedicated transformer accelerator and a QoS-guar…

View →
cs.NIEmpiricalRecentJul 24, 2026

CAPS: Fine-Tuning CCA Timing

Raphael Zailer, Isaac Keslassy

This paper proposes CAPS, a scheduling layer for data centers that separates rate computation and packet scheduling, reducing queue occupancy by up to 10x without throughput loss.

View →
cs.DCcs.AIcs.OSEmpiricalRecentJul 2, 2026

Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of Lines

Bojie Li

The paper proposes a method to improve the performance of fine-grained offloads on servers by overlapping the offload with other requests using server-side routing.

View →