20 results for “accelerator”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper proposes a heterogeneous neural network accelerator for multi-task RF signal recognition, achieving high accuracy and low latency for automatic modulation recognition, hardware-Trojan cover…
The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…
The paper introduces Chimera, a highly efficient and scalable MCU designed for ultra-low-power edge AI inference, achieving 3.1 TOPS/W by integrating a dedicated transformer accelerator and a QoS-guar…
The paper systematically characterizes column-level activation sparsity across various diffusion model architectures, demonstrating that element-level sparsity metrics significantly overestimate the a…
This paper investigates the thermal constraints of deploying AI compute infrastructure in space, comparing GPUs and compute-in-memory (CIM) accelerators using a co-design methodology.
This paper implements and measures the cost of running a streaming speech enhancer on a CPU, achieving a 3.3x speedup compared to the fp32 ONNX Runtime graph.
The paper introduces CA-AC-MPC, a CUDA-accelerated variant of Actor-Critic Model Predictive Control, which significantly reduces the training and inference latency of AC-MPC while maintaining state-of…
This paper presents Mana, a sim-to-real framework for dexterous articulated tool manipulation.
The paper proposes CLASP, an end-to-end system with IMC acceleration for continual learning on edge platforms, addressing challenges of noisy computation and poor support for resource-efficient traini…
Hawkeye is a system that allows perfect, precision-preserving reproduction of GPU-level matrix multiplication operations on a CPU, enabling efficient and trustworthy third-party auditing of machine le…
Siyuan Shen, Anton Korzh, John Bachan, Tiancheng Chen +9 more
This paper explores methods to reduce latency in GPU collective communications for large language model inference, achieving near-optimal designs with barrier-free synchronization and efficient use of…
The paper presents a field study on deploying large inference workloads on a non-GPU AI accelerator and identifies eight categories of limitations.