20 results for “Concept of Processing-near-Memory (PNM)”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper investigates the potential of real-world Processing-in-Memory (PIM) architectures, specifically using UPMEM, to accelerate cryptographic algorithms, demonstrating that distributing computat…
Bing Wu, Xueliang Wei, Shiyi Song, Yibo Liu +5 more
The paper introduces PATH, an in-situ indexing architecture for Processing-in-Memory systems that achieves higher throughput, lower tail latency, and fewer memory accesses than state-of-the-art scheme…
PIMID is an execution- and trace-driven full-system simulator for Processing-in-Memory systems, supporting multiple memory technologies, execution models, and placement of processing elements.
Sookyung Choi, Seungyong Lee, Kangkyu Park, Yunseo Chun +10 more
This paper presents NELSSA, a serving system that integrates GPUs with Processing-near-Memory (PNM) devices to efficiently handle mixed-length workloads in LLMs, achieving up to 5.5x decode throughput…
This paper demonstrates that using circuit-level compensation techniques and batch normalization recalibration can mitigate the impact of retention loss on inference accuracy degradation in a single-p…
This paper evaluates rank-order N-of-M encoding as an alternative to threshold-binary encoder in Sparse Distributed Memory systems and shows its outperformance in capacity experiments and robustness e…
SangHoon Cha, Jaewan Choi, Byeongho Kim, Yoonah Paik +2 more
This paper introduces a high-fidelity, integrated hardware-software simulator for LPDDR5X-PIM, enabling precise evaluation of system performance and energy efficiency.
This paper introduces phase-to-primitive decomposition to enable Monte Carlo tree search (MCTS) on in-memory computing (IMC) systems, achieving significant energy efficiency and performance improvemen…
This paper proposes a sparse cache with a novel allocation rule based on DP-means clustering for deep recurrent models, achieving full associative recall with only distinct items, outperforming other…
This paper provides the first systematic, isolated benchmarks of NIST-standardized post-quantum cryptography (ML-KEM and ML-DSA) on the highly constrained ARM Cortex-M0+ processor, showing performance…
The paper implements and evaluates several hardware priority queue architectures on modern FPGA platforms and provides a quantitative analysis.
The paper proposes a graph attention-based virtual metrology framework that accurately predicts film thickness in semiconductor deposition by modeling structured, directional dependencies among hetero…
Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin +5 more
Proposed EnerInfer framework manages energy efficiency, throughput, and thermal comfort for on-device LLM inference, improving energy efficiency up to 65% without QoE violation.
Shuo Huai, Hao Kong, Xiangzhong Luo, Shiqing Li +4 more
This paper addresses the obstacles of using Crossbar-based In-Memory Processing (IMP) accelerators for deep neural networks (DNNs) by reusing bit-shift units for multiplication, applying pruning metho…
This paper proposes Head-Chunked Multi-Stream Pipeline (HCMS) to exploit the computational independence of multi-head attention and achieve fine-grained communication-computation overlap, resulting in…
Jae Hyung Ju, Euijun Chung, Hritvik Taneja, Anish Saxena +3 more
This paper proposes TileLens, a system to mitigate read amplification in Large-Granularity Memory Systems (LGMS) for Large Language Model (LLM) inference by adopting a tile-major layout.
This paper characterizes the gap between current DNN-based speech enhancement systems and hearing aid constraints, and proposes a lightweight architecture to meet these constraints.