20 results for “neural network accelerator”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper investigates limitations of learning tanh neural networks under finite-precision computations and Lp accuracy guarantees.
This paper proposes Supervised Memory Training (SMT), a method for training nonlinear RNNs that sidesteps recurrent credit propagation entirely.
The elasticAI.explorer is an extensible, unified Python framework that simplifies hardware-aware Neural Architecture Search (NAS) by decoupling search space definition from model implementation and de…
This paper proposes a heterogeneous neural network accelerator for multi-task RF signal recognition, achieving high accuracy and low latency for automatic modulation recognition, hardware-Trojan cover…
OpenEye is a scalable, sparsity-aware FPGA-based hardware accelerator designed to efficiently execute common deep neural network operations, demonstrating favorable performance-resource trade-offs acr…
This paper introduces Mega, a digital architecture for Convolutional Spiking Neural Networks (SNNs) that addresses underutilization of parallelism and inflexibility in existing SNN accelerators throug…
Rania Zitouni, Nadine Bousdjira, Sarah Hasnaoui, Amel Sadoun +1 more
This paper compares and optimizes CUDA strategies for a shallow neural network, achieving a 1.41x speedup on a large dataset.
Shuo Huai, Hao Kong, Xiangzhong Luo, Shiqing Li +4 more
This paper addresses the obstacles of using Crossbar-based In-Memory Processing (IMP) accelerators for deep neural networks (DNNs) by reusing bit-shift units for multiplication, applying pruning metho…
The paper presents a field study on deploying large inference workloads on a non-GPU AI accelerator and identifies eight categories of limitations.
This paper proposes Hierarchical Block-Local Learning (HBLL), a framework for training deep neural networks without full end-to-end backpropagation, achieving $\mathcal{O}(\log N)$ parallel time compl…
Zheqi Shen, Jingbo Su, Zijin Wan, Yan Gu +1 more
ANNLib is a library for Approximate Nearest Neighbor Search (ANNS) providing high performance and flexible functionality using graph-based algorithms and data structures.
The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…
The paper introduces Automatically Differentiable Nonlinear Tensor Networks (ADNTNs) to achieve massive, structured compression of deep neural networks, demonstrating compression ratios up to 77,000x…
ACRONYM is a novel algorithm-hardware co-designed platform that enables high-recall, continuous approximate nearest neighbor search in memory for dynamic vector databases, achieving massive throughput…