~ similar to 2607.24196· 20 results
Bing Wu, Xueliang Wei, Shiyi Song, Yibo Liu +5 more
The paper introduces PATH, an in-situ indexing architecture for Processing-in-Memory systems that achieves higher throughput, lower tail latency, and fewer memory accesses than state-of-the-art scheme…
Yilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao +3 more
The paper introduces COSM, a cooperative scheduling framework to facilitate concurrent operation of Processing-in-Memory (PIM) and CPU tasks on mobile platforms, improving PIM throughput by up to 2.8x…
SangHoon Cha, Jaewan Choi, Byeongho Kim, Yoonah Paik +2 more
This paper introduces a high-fidelity, integrated hardware-software simulator for LPDDR5X-PIM, enabling precise evaluation of system performance and energy efficiency.
Haocong Luo, F. Nisa Bostancı, Ataberk Olgun, Maria Makeenkova +3 more
Ramulator 2.1 is an updated DRAM simulator with improved support for modern features, better usability, and comprehensive testing.
Valinor is a hardware-OS cooperative memory allocation substrate that introduces a programmable hardware allocation engine for improving performance and reducing energy consumption in virtual-to-physi…
This paper investigates the potential of real-world Processing-in-Memory (PIM) architectures, specifically using UPMEM, to accelerate cryptographic algorithms, demonstrating that distributing computat…
F. Nisa Bostanci, Haocong Luo, Ataberk Olgun, Maria Makeenkova +3 more
The authors of Ramulator 2.0 simulator challenge the claims made in a research paper about its performance and propose best practices to avoid simulator usage errors.
The paper introduces memorywire, a vendor-neutral JSON-Schema wire format and reference implementation designed to standardize and govern memory operations across disparate agent-memory frameworks.
Harshita Gupta, Mayank Kabra, Jaewoo Park, Priyam Mehta +8 more
The paper characterizes Homomorphic Encryption (HE) operations on a real-world Processing-In-Memory (PIM) system, demonstrating that while PIM is a viable alternative to CPUs/GPUs, performance is limi…
The paper introduces memorywire, a vendor-neutral JSON-Schema 2020-12 wire format and reference implementation to standardize and govern agent memory operations across diverse, proprietary agent-memor…
The paper characterizes 'dead-entry' TLB misses in GPUs, which occur when recently evicted translations are immediately re-walked, and proposes DEPOT, a Bloom filter mechanism that significantly reduc…
Ming-Yen Lee, Hanchen Yang, Faaiq Waqar, Harsono Simka +3 more
This paper investigates the energy efficiency of Large Language Model (LLM) serving using emerging memory technologies, specifically monolithic 3D (M3D) integration, and presents simulation results sh…
Yuchen Fan, Minghong Sun, Jikui Ma, Yunpeng Xu +18 more
The paper presents CHASE, an application-driven framework that explores physically feasible Cross-layer Heterogeneous System architectures for executing workloads with diverse requirements.
The I/O Resource Manager (IORM) is presented as a multi-stage distributed scheduler to maintain consistent performance and enforce global I/O limits in shared-nothing disaggregated storage clusters.
This paper argues that idle cores on contemporary multicore processors can return compute capacity and proposes a budget-centric view of service core systems.
Sookyung Choi, Seungyong Lee, Kangkyu Park, Yunseo Chun +10 more
This paper presents NELSSA, a serving system that integrates GPUs with Processing-near-Memory (PNM) devices to efficiently handle mixed-length workloads in LLMs, achieving up to 5.5x decode throughput…
Adithya Kumar, Aditya Basu, Jacob Kahn, Parth Malani +2 more
The paper introduces PRISM, an evaluation framework for assessing POSIX storage systems for AI research workloads, and compares Lustre and NFS based systems.
Hai Duc Nguyen, Tekin Bicer, Kyle Chard, Ian Foster +1 more
Researchers used a large language model to generate checkpoint/restart code for MPI applications, achieving comparable efficiency to human-engineered solutions.
The paper proposes the Intelligent Computing Architecture Model (ICAM), a six-layer framework that unifies disparate concepts in model-native computing by viewing the LLM stack through a dual-plane ar…
Zhuoren Ye, Tianyu Wo, Dinghao Xue, Mingming Zhang +3 more
This paper proposes CrossPool, a serving engine for cold Machine Learning Models (LLMs) that separates weights and KV-cache into two GPU memory pools to improve GPU memory utilization and long-context…