ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Concept of Processing-near-Memory (PNM)”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.CRcs.ARcs.DCRecentMay 19, 2026

Taking Cryptography Out of the Data Path via Near-Memory Processing in DRAM

Nicola Barcarolo, Brahmaiah Gandham, Mohammad Sadrosadati, Roberto Passerone +2 more

This paper investigates the potential of real-world Processing-in-Memory (PIM) architectures, specifically using UPMEM, to accelerate cryptographic algorithms, demonstrating that distributing computat…

View →
cs.ARcs.ETEmpiricalRecentJun 30, 2026

In-situ Indexing via Memristive Content-Addressable Memory

Bing Wu, Xueliang Wei, Shiyi Song, Yibo Liu +5 more

The paper introduces PATH, an in-situ indexing architecture for Processing-in-Memory systems that achieves higher throughput, lower tail latency, and fewer memory accesses than state-of-the-art scheme…

View →
cs.AREmpiricalRecentJul 27, 2026

PIMID: A Full-System Simulator with Intricacy and Diversity for Processing-in-Memory

Yuan He, Masaaki Kondo, Galen M. Shipman, Jered B. Dominguez-Trujillo +2 more

PIMID is an execution- and trace-driven full-system simulator for Processing-in-Memory systems, supporting multiple memory technologies, execution models, and placement of processing elements.

View →
cs.ARNEWEmpiricalJul 29, 2026

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement

Sookyung Choi, Seungyong Lee, Kangkyu Park, Yunseo Chun +10 more

This paper presents NELSSA, a serving system that integrates GPUs with Processing-near-Memory (PNM) devices to efficiently handle mixed-length workloads in LLMs, achieving up to 5.5x decode throughput…

View →
cs.ARcs.NEphysics.app-phEmpiricalRecentJul 27, 2026

Mitigating the Impact of Retention Loss on Inference Accuracy in 65 nm Single-Poly Floating-Gate Analog In-Memory Computing

Mirko Brazzini, Giulio Filippeschi, Alessandro Catania, Sebastiano Strangio +1 more

This paper demonstrates that using circuit-level compensation techniques and batch normalization recalibration can mitigate the impact of retention loss on inference accuracy degradation in a single-p…

View →
cs.LGcs.NEEmpiricalRecentJul 3, 2026

Rank-Order N-of-M Codes for Sparse Distributed Memory: Disentangling Representation and Learning Effects in Noise Robustness Against Contemporary Neuromorphic Architectures

Joy Bose

This paper evaluates rank-order N-of-M encoding as an alternative to threshold-binary encoder in Sparse Distributed Memory systems and shows its outperformance in capacity experiments and robustness e…

View →
cs.ARcs.AIRecentMay 30, 2026

LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM

SangHoon Cha, Jaewan Choi, Byeongho Kim, Yoonah Paik +2 more

This paper introduces a high-fidelity, integrated hardware-software simulator for LPDDR5X-PIM, enabling precise evaluation of system performance and energy efficiency.

View →
cs.ARcs.AIcs.ETEmpiricalRecentJul 24, 2026

Multi-primitive in-memory computing for Monte Carlo tree search

Tergel Molom-Ochir, Benjamin F. Morris, Yintao He, Archit Gajjar +5 more

This paper introduces phase-to-primitive decomposition to enable Monte Carlo tree search (MCTS) on in-memory computing (IMC) systems, achieving significant energy efficiency and performance improvemen…

View →
cs.LGcs.AIcs.CLEmpiricalRecentJul 10, 2026

Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention

Siddharth Pal, Viktoria Rojkova

This paper proposes a sparse cache with a novel allocation rule based on DP-means clustering for deep recurrent models, achieving full associative recall with only distinct items, outperforming other…

View →
cs.CRcs.ARcs.PFRecentMar 19, 2026

Benchmarking NIST-Standardised ML-KEM and ML-DSA on ARM Cortex-M0+: Performance, Memory, and Energy on the RP2040

Rojin Chhetri

This paper provides the first systematic, isolated benchmarks of NIST-standardized post-quantum cryptography (ML-KEM and ML-DSA) on the highly constrained ARM Cortex-M0+ processor, showing performance…

View →
cs.AREmpiricalRecentJul 22, 2026

Revisiting Hardware Priority Queue Architectures

Qihang Wu, Austin Rovinski

The paper implements and evaluates several hardware priority queue architectures on modern FPGA platforms and provides a quantitative analysis.

View →
cs.CERecentMay 30, 2026

Graph Attention-Based Virtual Metrology for Film Deposition Processes in Semiconductor Manufacturing

Tao Han, Suk Ki Lee, Hyunwoong Ko

The paper proposes a graph attention-based virtual metrology framework that accurately predicts film thickness in semiconductor deposition by modeling structured, directional dependencies among hetero…

View →
cs.SEcs.LGcs.OSEmpiricalRecentJun 22, 2026

EnerInfer: Energy-Aware On-Device LLM Inference

Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin +5 more

Proposed EnerInfer framework manages energy efficiency, throughput, and thermal comfort for on-device LLM inference, improving energy efficiency up to 65% without QoE violation.

View →
cs.AREmpiricalRecentJul 9, 2026

CRIMP: Compact & Reliable DNN Inference on In-Memory Processing via Crossbar-Aligned Compression and Non-ideality Adaptation

Shuo Huai, Hao Kong, Xiangzhong Luo, Shiqing Li +4 more

This paper addresses the obstacles of using Crossbar-based In-Memory Processing (IMP) accelerators for deep neural networks (DNNs) by reusing bit-shift units for multiplication, applying pruning metho…

View →
cs.DCEmpiricalRecentJul 2, 2026

HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention

Chao Yuan, Pan Li, Yingnan Sun, Jing Liu

This paper proposes Head-Chunked Multi-Stream Pipeline (HCMS) to exploit the computational independence of multi-head attention and achieve fine-grained communication-computation overlap, resulting in…

View →
cs.AREmpiricalRecentJul 13, 2026

Reliable Associative Lookup in Content-Addressable Memory

Fan Li, Yanan Guo, Xin Xin

This paper introduces a new protection code design for Content Addressable Memory (CAM) to ensure reliability.

View →
cs.AREmpiricalRecentJul 4, 2026

TileLens: Efficiently Using Large-Granularity Memory Systems with Transparent Two-Dimensional Memory Layout

Jae Hyung Ju, Euijun Chung, Hritvik Taneja, Anish Saxena +3 more

This paper proposes TileLens, a system to mitigate read amplification in Large-Granularity Memory Systems (LGMS) for Large Language Model (LLM) inference by adopting a tile-major layout.

View →
cs.SDcs.AReess.ASRecentJun 2, 2026

Feasibility of Time-Domain DNN-Based Speech Enhancement on Embedded FPGA for Hearing Aid

Feyisayo Olalere, Umut Altin, Kiki van der Heijden, Marcel van Gerven

This paper characterizes the gap between current DNN-based speech enhancement systems and hearing aid constraints, and proposes a lightweight architecture to meet these constraints.

View →